Automatic Text Categorization In Terms Of Genre And Author

نویسندگان

  • Efstathios Stamatatos
  • George K. Kokkinakis
  • Nikos D. Fakotakis
چکیده

The two main factors that characterize a text are its content and its style. Both of them can be used as categorization means. In this paper we present an approach to text categorization in terms of genre and author for Modern Greek. In contrast to hitherto stylometric approaches, we attempt to take full advantage of existing natural language processing (NLP) tools. To this end, we propose a set of style markers including analysislevel measures that represent the way in which the input text has been analyzed and capture useful stylistic information without additional cost. We present a set of smallscale but reasonable experiments in text genre detection, author identification as well as author verification tasks and show that the performance of the proposed method is better in comparison with the most popular distributional lexical measures, i.e., functions of vocabulary richness and frequencies of occurrence of the most frequent words. All the presented experiments are based on unrestricted text downloaded from the World Wide Web (WWW) without any manual text preprocessing or text sampling. Various performance issues regarding the training set size and the significance of the proposed Stamatatos et al. Text Categorization 2 style markers are discussed. Our system can be used in any application that requires fast and easily adaptable text categorization in terms of stylistically homogeneous categories. Moreover, the procedure of defining analysis-level markers can be followed in order to employ already existing text processing tools for the extraction of useful stylistic information.

برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

مدل دو مرحله ای شکاف- گلچین برای نمایه سازی خودکار متون فارسی

Purpose: Each language has its own problems. This leads to consider appropriate models for automatic indexing of every language. These models should concern the exhaustificity and specificity of indexing.   This paper aims at introduction and evaluation of a model which is suited for Persian automatic indexing. This model suggests to break the text into the particles of candidate terms and to c...

متن کامل

A survey on Automatic Text Summarization

Text summarization endeavors to produce a summary version of a text, while maintaining the original ideas. The textual content on the web, in particular, is growing at an exponential rate. The ability to decipher through such massive amount of data, in order to extract the useful information, is a major undertaking and requires an automatic mechanism to aid with the extant repository of informa...

متن کامل

A Genre Analysis of Persian Research Article Abstracts: Communicative Moves and Author Identity

Most studies within the area of genre analysis, particularly those conducted in Iran, have exclusively used text analysis. While such investigations have led to important understandings of generic features of texts, it can be argued that incorporating interview data for triangulation can lead to better understanding of generic features of texts. Along this line, this paper reports the results o...

متن کامل

Using syntactic features to predict author personality from text

The style in which a text is written re ects an array of meta-information concerning the text (e.g., topic, register, genre) and its author (e.g., gender, region, age, personality). The eld of stylometry addresses these aspects of style. A successful methodology, borrowed from text categorisation research, takes a two-stage approach which (i) achieves automatic selection of features with high p...

متن کامل

Genre Categorization and Modeling for Broadcast Speech Transcription

Broadcast News (BN) speech recognition transcription has attracted research due to the challenges of the task since the mid 1990’s. More recently, research has been moving towards more spontaneous broadcast data, commonly called Broadcast Conversation (BC) speech. Considering the large style difference between BN and BC genres, specific modeling of genres should intuitively result in improved s...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 2000